Classical Recommender Systems
Recommender systems dictate modern internet consumption—they power YouTube homepages, Amazon product suggestions, and TikTok feeds.
The Two Primary Paradigms
- Content-Based Filtering: Recommends items based on item features. (e.g., If you watch an action movie starring Tom Cruise, the system recommends another action movie starring Tom Cruise).
- Collaborative Filtering: Recommends items based on user behavior and interactions. (e.g., "Users who bought this also bought...").
Matrix Factorization
The core engine of classical Collaborative Filtering is Matrix Factorization (which won the famous $1M Netflix Prize).
Imagine a massive, sparse matrix where Rows are Users, Columns are Movies, and the cells are Ratings (1-5 stars). Most cells are empty. Matrix Factorization uses linear algebra (specifically SVD - Singular Value Decomposition) to split this giant matrix into two smaller, dense matrices:
- A User Embedding matrix.
- An Item Embedding matrix.
When you take the dot product of a User Embedding and an Item Embedding, you get the predicted rating that the user would give that movie!
Python Implementation: Matrix Factorization
import numpy as np
from sklearn.decomposition import TruncatedSVD
# Simulated User-Item Rating Matrix (Users x Items)
# 0 means unrated.
ratings_matrix = np.array([
[5, 3, 0, 1],
[4, 0, 0, 1],
[1, 1, 0, 5],
[1, 0, 0, 4],
[0, 1, 5, 4],
])
# Perform Singular Value Decomposition (Matrix Factorization)
# We compress the items into 2 latent features
svd = TruncatedSVD(n_components=2)
user_embeddings = svd.fit_transform(ratings_matrix)
item_embeddings = svd.components_
# Reconstruct the matrix to see the predicted ratings for unrated items!
predicted_ratings = np.dot(user_embeddings, item_embeddings)
print("Original Matrix:\n", ratings_matrix)
print("\nPredicted Ratings Matrix:\n", np.round(predicted_ratings, 1))